Papers with Language Agents

4 papers
Rethinking Stateful Tool Use in Multi-Turn Dialogues: Benchmarks and Challenges (2025.findings-acl)

Copied to clipboard

Challenge: Existing benchmarks that assess Language Models (LMs) as Language Agents (LAs) for tool use focus on stateless, single-turn interactions or partial evaluations, overlooking the inherent stateful nature of interactions in multi-turn applications.
Approach: They propose a multi-turn dialogue dataset with stateful tool interactions considering the whole life cycle of tool use across six key tasks in three stages . they also build VirtualMobile – an embodied virtual mobile evaluation environment to simulate API calls and assess the robustness of the created APIs.
Outcome: The proposed dataset evaluates 13 open- and closed-source LLMs and provides detailed analysis at each stage.
Towards Uncertainty-Aware Language Agent (2024.findings-acl)

Copied to clipboard

Challenge: Existing Language Agents neglect the notion of uncertainty during interactions with external worlds.
Approach: They propose a framework that orchestrates the interaction between the agent and the external world using uncertainty quantification.
Outcome: The proposed framework improves performance on 3 representative tasks and lowers reliance on external world.
MetaReflection: Learning Instructions for Language Agents using Past Reflections (2024.emnlp-main)

Copied to clipboard

Challenge: Large Language Models (LLMs) have gained popularity due to their ability to generate human-like text and solve complex tasks.
Approach: They propose an offline reinforcement learning technique that augments a semantic memory based on experiential learnings from past trials.
Outcome: The proposed technique boosts Language agents’ performance by 4 % to 16.82 % over the raw GPT-4 baseline and performs on par with existing state-of-the-art prompt optimization techniques while requiring fewer LLM calls.
Can a Single Model Master Both Multi-turn Conversations and Tool Use? CoALM: A Unified Conversational Agentic Language Model (2025.acl-long)

Copied to clipboard

Challenge: Large Language Models (LLMs) with API-calling capabilities enabled building effective Language Agents (LA) current approaches excel in one domain but underperform in the other.
Approach: They propose a unified approach that integrates both conversational and agentic capabilities.
Outcome: The proposed model outperforms top domain-specific models across three benchmarks.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations